Visual question answering from another perspective: CLEVR mental rotation tests

نویسندگان

چکیده

Different types of mental rotation tests have been used extensively in psychology to understand human visual reasoning and perception. Understanding what an object or scene would look like from another viewpoint is a challenging problem that made even harder if it must be performed single image. We explore controlled setting whereby questions are posed about the properties was observed viewpoint. To do this we created new version CLEVR dataset call Mental Rotation Tests (CLEVR-MRT). Using CLEVR-MRT examine standard methods, show how they fall short, then novel neural architectures involve inferring volumetric representations scene. These volumes can manipulated via camera-conditioned transformations answer question. efficacy different model variants through rigorous ablations demonstrate representations.

برای دانلود باید عضویت طلایی داشته باشید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

From Question Answering to Visual Exploration

Success in Question Answering has been traditionally measured by precision and recall, which are good metrics for identifying specific best answer(s) that might be obtained by a lookup type of search. These metrics do not address the many information gathering techniques in exploratory interactions. In this paper, we present an integrated Question Answering environment that combines a visual an...

متن کامل

Investigating Embedded Question Reuse in Question Answering

The investigation presented in this paper is a novel method in question answering (QA) that enables a QA system to gain performance through reuse of information in the answer to one question to answer another related question. Our analysis shows that a pair of question in a general open domain QA can have embedding relation through their mentions of noun phrase expressions. We present methods f...

متن کامل

Revisiting Visual Question Answering Baselines

Visual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding. Many of the recently proposed VQA systems include attention or memory mechanisms designed to support “reasoning”. For multiple-choice VQA, nearly all of these systems train a multi-class classifier on image and question features to predict ...

متن کامل

iVQA: Inverse Visual Question Answering

In recent years, visual question answering (VQA) has become topical as a long-term goal to drive computer vision and multi-disciplinary AI research. The premise of VQA’s significance, is that both the image and textual question need to be well understood and mutually grounded in order to infer the correct answer. However, current VQA models perhaps ‘understand’ less than initially hoped, and in...

متن کامل

Speech-Based Visual Question Answering

This paper introduces the task of speech-based visual question answering (VQA), that is, to generate an answer given an image and an associated spoken question. Our work is the first study of speechbased VQA with the intention of providing insights for applications such as speech-based virtual assistants. Two methods are studied: an end to end, deep neural network that directly uses audio wavef...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: Pattern Recognition

سال: 2023

ISSN: ['1873-5142', '0031-3203']

DOI: https://doi.org/10.1016/j.patcog.2022.109209